Goto

Collaborating Authors

 Mexico City


La Catalina ends Flammer's 1,128-day reign as AAA Reina de Reinas champion at Triplemania 34

FOX News

WWE star Stephanie Vaquer captures Women's World Championship at Chile live event in surprising moment Ben Shelton loses in US Open final to Zverev, but America has found its next big men's tennis star Jeremiyah Love makes immediate statement with touchdown on first career drive in Cardinals' upset win Josh Allen torches one of NFL's best defenses to make case it's time to move on from 2025 playoff loss Breanna Stewart's double-double, Caitlin Clark's emergence off bench help US to World Cup win over France USC fans trade punches with one another in the stands because winning sometimes isn't enough FIBA Secretary General Andreas Zagklis speaks out on Caitlin Clark's impact on growth of World Cup Prosecution in Lindsay Clancy trial may have'alienated' jurors, criminal defense attorney says You treat a nuclear power'significantly different' than a non-nuclear power: Gen. Keith Kellogg Accepting political violence as a form of expression is'dangerous,' Jonathan Turley warns'The Squad' faces backlash over claims'modern-day lynchings' are now common in America There is no'moderate socialism,' Republican consultant says Prince Harry's reported Princess Diana movie project sparks controversy Prince Harry's reported Princess Diana movie project sparks controversy La Catalina made her AAA debut in April to the surprise of Mexican fans and immediately worked her way into the Reina de Reinas Championship picture. She won a gauntlet match to get a title shot at Triplemania 34, and on Sunday, she ended the long reign of Flammer to win the title. La Catalina faces Flammer during Triplemania 34: Night 2 at Arena Ciudad de Mexico in Mexico City, Mexico, on Sept. 13, 2026. La Catalina had Flammer on the ropes late in the match. She went for a diving splash but didn't get all of it as Flammer put her knees up at the last second.


The 10 longest airplane flights on Earth

Popular Science

These ultra-long commercial flights are becoming more common thanks to design innovations and fuel capacity. More information Adding us as a Preferred Source in Google by using this link indicates that you would like to see more of our content in Google News results. At 24 hours and 24 minutes of flight time, Project Sunrise broke the old record for longest-ever flight by a commercial-style aircraft. Breakthroughs, discoveries, and DIY tips sent six days a week. By signing up, you confirm you are 16+, will receive newsletters and promotional content and agree to our Terms of Use and acknowledge the data practices in our Privacy Policy .


European Lawmakers Demand Investigation Into FIFA President Over U.S. Red Card Reversal

TIME - Tech

Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW?


Neither Valid nor Reliable Investigating the Use of LLMs as Judges

Neural Information Processing Systems

Evaluating natural language generation (NLG) systems remains a core challenge of natural language processing (NLP), further complicated by the rise of large language models (LLMs) that aim to be general-purpose. Recently, large language models as judges (LLJs) have emerged as a promising alternative to traditional metrics, but their validity remains underexplored. This position paper argues that the current enthusiasm around LLJs may be premature, as their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. Drawing on measurement theory from the social sciences, we identify and critically assess four core assumptions underlying the use of LLJs: their ability to act as proxies for human judgment, their capabilities as evaluators, their scalability, and their cost-effectiveness. We examine how each of these assumptions may be challenged by the inherent limitations of LLMs, LLJs, or current practices in NLG evaluation. To ground our analysis, we explore three applications of LLJs: text summarization, data annotation, and safety alignment. Finally, we highlight the need for more responsible evaluation practices in LLJs evaluation, to ensure that their growing role in the field supports, rather than undermines, progress in NLG.


Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Neural Information Processing Systems

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as'safety' and'robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.


How Mexican World Cup Stadiums Achieved FIFA's Environmental Certifications

WIRED

Venues hosting the 2026 World Cup must meet high standards to obtain environmental certifications, but FIFA also requires that they use natural grass, which is water-intensive to maintain. Estadio Banorte, formerly called Azteca stadium, in Mexico City. Because of their scale, soccer stadiums require a fair amount of energy and water. In that time, they also generate large volumes of waste, mainly plastics and food trash. For the 2026 World Cup, the first to be held in three countries in 16 different stadiums, FIFA maintained the requirement that the venues must have LEED environmental certifications, which measure performance in water, energy, and waste management.


Diana Flores

TIME - Tech

Follow this author to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Flag football will step firmly onto the global stage at the 2028 Olympics in Los Angeles, and Diana Flores is one of the key figures who helped it get there. The Mexico City native began playing at age 8 and made her country's national team by 16.


I own 20 axolotls - people need to know they're not easy to look after

BBC News

I own 20 axolotls - people need to know they're not easy to look after When Emma Honeyfield's daughter Amber asked for an axolotl for her birthday, Emma never imagined it would lead to a collection of 20. The 37-year-old bought her daughter's first axolotl, Stitch, in September and has since fallen in love with their calming nature. Emma said Amber, eight, had always been difficult to buy for, so when she asked for one for her birthday, she couldn't say no. And the family, from Tredegar, Blaenau Gwent, are far from alone in seeking out the amphibians, which are critically endangered and only found in lakes and wetlands in southern Mexico City . The animal's cute, smiling face and appearance in the hugely popular Minecraft and Roblox games has seen an increase in the number of people keeping them as pets.


A Bayesian Perspective on the Role of Epistemic Uncertainty for Delayed Generalization in In-Context Learning

arXiv.org Machine Learning

In-context learning enables transformers to adapt to new tasks from a few examples at inference time, while grokking highlights that this generalization can emerge abruptly only after prolonged training. We study task generalization and grokking in in-context learning using a Bayesian perspective, asking what enables the delayed transition from memorization to generalization. Concretely, we consider modular arithmetic tasks in which a transformer must infer a latent linear function solely from in-context examples and analyze how predictive uncertainty evolves during training. We combine approximate Bayesian techniques to estimate the posterior distribution and we study how uncertainty behaves across training and under changes in task diversity, context length, and context noise. We find that epistemic uncertainty collapses sharply when the model groks, making uncertainty a practical label-free diagnostic of generalization in transformers. Additionally, we provide theoretical support with a simplified Bayesian linear model, showing that asymptotically both delayed generalization and uncertainty peaks arise from the same underlying spectral mechanism, which links grokking time to uncertainty dynamics.


Mexico City's 'Xoli' Chatbot Will Help World Cup Tourists Navigate the City

WIRED

The launch of "Xoli" adds to the technological efforts promoted by the federal government to turn the 2026 World Cup into an engine of development for the entire country. Xoli, the new chatbot, is named after the axolotl, a salamander with external gills. The Government of Mexico City has launched Xoli, a chatbot that will provide information on services, tourism, and cultural offerings. The platform was designed to meet the demand of the millions of visitors expected to arrive during the 2026 FIFA World Cup . However, the authorities assure that the tool will remain active once the sporting event is over, with the aim of promoting economic activities and facilitating access to public services in the capital.